Papers with RL methods
Learning Natural Language Generation with Truncated Reinforcement Learning (2022.naacl-main)
Copied to clipboard
| Challenge: | Existing approaches to train conditional languagemodels without supervised learning fail to scale to large action spaces, thus allowing to train a language agent by only interacting with its environment without any task-specific prior knowledge. |
| Approach: | They propose an original approach to train conditional languagemodels without supervised learning by only using reinforcement learning. |
| Outcome: | The proposed approach avoids the dependency to labelled datasets and reduces pretrained policy flaws such as language or exposure biases. |
Fine-Grained Reward Optimization for Machine Translation using Error Severity Mappings (2026.tacl-1)
Copied to clipboard
Miguel Moura Ramos, Tomás Almeida, Daniel Vareta, Filipe Azevedo, Sweta Agrawal, Patrick Fernandes, André F. T. Martins
| Challenge: | Reinforcement learning (RL) is an effective and robust method for training neural machine translation systems. |
| Approach: | They propose a method that leverages fine-grained, token-level quality assessments . they use a state-of-the-art quality estimation system as their token- level reward model . |
| Outcome: | The proposed approach leverages fine-grained, token-level quality assessments along with error severity levels to improve translation quality. |
Reinforcement Learning with Token-level Feedback for Controllable Text Generation (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for controllable text generation are guided by coarse-grained feedback, which may lead to suboptimal performance owing to semantic twists or progressions within sentences. |
| Approach: | They propose a reinforcement learning algorithm which formulates TOken-LEvel rewards for controllable text generation and employs a "first-quantize-then-noise" paradigm to enhance the robustness of the RL algorithm. |
| Outcome: | The proposed algorithm can achieve superior performance on single-attribute and multi-attract control tasks. |
Learning from Mistakes: Iterative Prompt Relabeling for Text-to-Image Diffusion Model Training (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in diffusion models have shown impressive performance in many domains, but their ability to follow instructions is still unsatisfactory. |
| Approach: | They propose an algorithm that aligns images to text through iterative image sampling and prompt relabeling with feedback. |
| Outcome: | The proposed algorithm improves on the spatial relation VISOR benchmark by 15.22% compared to previous methods. |
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning (2026.findings-acl)
Copied to clipboard
Jinyang Wu, Chonghua Liao, Mingkuan Feng, Shuai Zhang, Zhengqi Wen, Haoran Luo, Ling Yang, Huazhe Xu, Jianhua Tao
| Challenge: | Existing RL methods rely on unstructured self-sampling to fit scalar rewards, resulting in inefficient rollouts. |
| Approach: | They propose a structured template-guided RL framework that augments policy optimization with explicit template guidance. |
| Outcome: | Experiments show that TemplateRL outperforms GRPO and GRPI by 99% on AIME and 41% on AMC with superior stability on weak models and remarkable cross-domain generalization. |
LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization (2025.emnlp-main)
Copied to clipboard
Qi Zhang, Shouqing Yang, Lirong Gao, Hao Chen, Xiaomeng Hu, Jinglei Chen, Jiexiang Wang, Sheng Guo, Bo Zheng, Haobo Wang, Junbo Zhao
| Challenge: | Recent research focuses on integrating reasoning capabilities into the realm of retrieval-augmented generation (RAG) via outcome-supervised reinforcement learning (RL). |
| Approach: | They propose a process-level reward module to mitigate the unawareness of intermediate reasoning steps in outcome-level supervision without additional annotation. |
| Outcome: | The proposed framework can boost LLMs’ reasoning ability by integrating external knowledge sources through retrieval-augmented generation (RAG) The proposed model can mitigate the unawareness of intermediate reasoning steps in outcome-level supervision without additional annotation. |
Distill and Align Decomposition for Enhanced Claim Verification (2026.findings-eacl)
Copied to clipboard
Jabez Magomere, Elena Kochkina, Samuel Mensah, Simerjot Kaur, Fernando Acero, Arturo Oncevay, Charese Smiley, Xiaomo Liu, Manuela Veloso
| Challenge: | Existing methods for complex claim verification struggle to align decomposition quality with verification performance. |
| Approach: | They propose a reinforcement learning approach that optimizes decomposition quality and verifier alignment using Group Relative Policy Optimization. |
| Outcome: | The proposed method outperforms prompt-based approaches and existing methods in six evaluation settings. |
Improving Large Language Models via Fine-grained Reinforcement Learning with Minimum Editing Constraint (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing reinforcement learning methods do not provide fine-grained supervision for complex reasoning tasks. |
| Approach: | They propose a reinforcement learning method that incorporates a generative model as the reward model and a token-level supervision model for RL training. |
| Outcome: | Experiments on 8 tasks show the proposed method is effective . |
A Teacher-Student Framework for Maintainable Dialog Manager (D18-1)
Copied to clipboard
| Challenge: | Reinforcement learning (RL) is an attractive solution for task-oriented dialog systems . but extending RL-based systems to handle new intents and slots requires a system redesign . |
| Approach: | They propose a teacher-student framework to extend RL-based dialog systems . they propose to specify constraints held in the new dialog manager . |
| Outcome: | The proposed framework makes no assumption about unsupported intents and slots, making it possible to improve RL-based systems incrementally. |
LLM-Coordination: Evaluating and Analyzing Multi-agent Coordination Abilities in Large Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated emergent common-sense reasoning and Theory of Mind (ToM) capabilities, making them promising candidates for developing coordination agents. |
| Approach: | They propose to use Large Language Models (LLMs) to analyze coordination models in Pure Coordination settings where agents must cooperate to maximize gains. |
| Outcome: | The proposed benchmark evaluates LLMs through two distinct tasks: Agentic Coordination and Coordination Question Answering. |
Rewarding What Matters: Step-by-Step Reinforcement Learning for Task-Oriented Dialogue (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing RL methods focus on generation tasks while neglecting dialogue state tracking (DST) for understanding. |
| Approach: | They propose a method that integrates RL into both understanding and generation tasks by introducing step-by-step rewards throughout the token generation. |
| Outcome: | The proposed approach achieves state-of-the-art results on three widely used datasets. |
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) relies on complex methodologies like Proximal Policy Optimization (PPO) that require extensive hyper-parameter tuning and pose challenges in sample efficiency and stability. |
| Approach: | They propose an innovative framework that leverages direct preference optimization techniques but extends them by estimating the conditionally optimal policy directly from the model’s responses. |
| Outcome: | The proposed framework matches and exceeds the effectiveness of Proximal Policy Optimization (PPO) in terms of convergence speed and alignment of model responses with human preferences. |
Efficient (Soft) Q-Learning for Text Generation with Limited Good Data (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Maximum likelihood estimation (MLE) is the predominant method for training text generation models. |
| Approach: | They propose a new RL formulation for text generation from the soft Q-learning perspective using path consistency learning to combine the best of on-/off-policy updates and learn effectively from sparse reward. |
| Outcome: | The proposed approach outperforms MLE and previous RL methods in a wide range of tasks. |
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for reinforcement learning for large language models do not accurately assess generalization. |
| Approach: | They propose three core principles for designing more faithful benchmarks: sufficient difficulty, balanced evaluation, and distributional robustness. |
| Outcome: | The proposed benchmarks do not accurately assess generalization across distribution shifts, difficulty levels, and counterfactual scenarios. |
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization (2026.acl-long)
Copied to clipboard
Yifeng Ding, Hung Le, Songyang Han, Kangrui Ruan, Zhenghui Jin, Varun Kumar, Zijian Wang, Anoop Deoras
| Challenge: | Current reinforcement learning methods suffer from coarse-grained, trajectory-level rewards that provide insufficient learning signals for complex multi-turn interactions, leading to training stagnation. |
| Approach: | They propose a novel RL algorithm for training large language models for multi-turn tool-integrated reasoning (TIR) that incorporates three innovations: turn-level reward assignment that provides fine-grained feedback for individual turns, return-based advantage estimation where normalized discounted returns are calculated as advantages, and self-supervised reward shaping that exploits self-supervision signals from generated code to densify sparse binary outcome-based rewards. |
| Outcome: | The proposed algorithm outperforms GRPO by 3.0% across diverse math reasoning benchmarks and improves grepo by 3.9% on commonsense reasoning and program synthesis tasks. |